Back

BMC Medical Education

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match BMC Medical Education's content profile, based on 21 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.

1
Resident Physician Selection Practices and Professionalism-Related Difficulties in Japan: A Nationwide Cross-sectional Survey

Sekine, M.; Nishizaki, Y.; Watari, T.; Shikino, K.; Fukui, S.; Nagasaki, K.; Nojima, M.; Shimizu, T.; Yamamoto, Y.; Kobayashi, H.; Tokuda, Y.

2026-08-07 medical education 10.64898/2026.08.05.26359757 medRxiv
Top 0.1%
64.0%
Show abstract

Introduction: Postgraduate clinical training is crucial for developing professional competence, communication skills, and effective teamwork. Although resident physician selection is crucial, little is known about how Japanese residency programs select residents and whether selection practices are associated with difficulties during training. Methods: We conducted a nationwide cross-sectional survey of residency programs participating in Japan's 2023 General Medicine In-Training Examination (GM-ITE). Program directors completed a questionnaire assessing selection methods, interview content, quality-assurance measures, and resident difficulties, defined as at least one postgraduate year 1 or 2 resident physician receiving disciplinary action or a severe warning. Free-text responses were coded using the Situation, Task, Action, and Result framework. Associations between selection methods and resident difficulties were examined using adjusted logistic regression models controlling for hospital type and number of GM-ITE examinees. Results: Of 151 participating physician-selection programs, 150 provided valid responses. Interviews were used by 90.1% of programs and were identified as the most important selection component by 87.3%. Thirty-five programs (23.3%) reported difficulties with resident physicians, involving professionalism and workplace conduct including rule, ethics, or boundary violations, work avoidance or unavailability, and inappropriate communication. Use of applicants' pre-clinical-clerkship computer-based test scores as a selection criterion was associated with resident difficulties (adjusted OR, 4.60; 95% CI, 1.50-14.11; P = 0.008; FDR-adjusted P = 0.048). No significant associations were observed for essays, academic tests, medical school grades, or personality assessment. Program-level GM-ITE total and domain scores did not differ significantly between programs with and without reported resident physician difficulties. Discussion: Resident physician selection in Japan is highly interview-centered; reported difficulties were more often related to professionalism and workplace conduct than to knowledge deficits. Although these exploratory findings are program-level, they highlight the importance of strengthening the quality assurance processes in resident selection systems, particularly for assessing professionalism-related attributes in applicants.

2
Comparative evaluation of case studies and role-playing techniques in enhancing problem-based learning outcomes in Medical Biochemistry education- An educational randomized controlled trial.

Kadam, C. Y.; Dhok, A.; Khanna, K.; Waghmare, T.

2026-07-29 medical education 10.64898/2026.07.27.26359062 medRxiv
Top 0.1%
63.7%
Show abstract

Background and objective: Problem-based learning (PBL) strategies such as case studies and role-playing are increasingly adopted in medical education to promote critical thinking, communication, and teamwork. However, comparative evidence regarding their effectiveness in medical biochemistry remains limited. This study aims to evaluate and compare the impact of case studies and role-playing on learning outcomes, engagement and skill development among first year MBBS students. Material and methods: An educational randomized controlled trial (RCT) with a within-subjects crossover design was conducted in the Department of Biochemistry involving forty-five first-year MBBS students. Participants were randomly assigned to receive case studies and role-playing in a counterbalanced sequence. Learning outcomes were assessed using pre- and post- tests, a structured feedback questionnaire and qualitative reflections. Data were analysed using appropriate parametric and non-parametric tests, along with adjusted analysis using ANCOVA to account for baseline performance. Results: All participants completed both methods, with no significant difference in baseline knowledge between arms (P=0.12). Both case studies and role-playing produced significant pre- to post- test improvements (P<0.001). Unadjusted analyses showed higher learning gains with role-playing, but pre-crossover ANCOVA demonstrated comparable effectiveness after baseline adjustment. Exploratory pooled analysis suggested a moderate cumulative advantage for role-playing. Student feedback indicated higher engagement and perceived skill development with role-playing, while both methods were viewed as relevant and educationally valuable. Conclusion: Both case studies and role-playing were effective PBL strategies for improving understanding and application of medical biochemistry concepts. Although student feedback and unadjusted analyses indicated higher engagement with role-playing, adjusted analysis at initial exposure showed comparable effectiveness between the two methods. Role-playing supported interactive learning and AETCOM -related skills, while case studies strengthened structured analytical reasoning. A combined approach may therefore offer a balanced and comprehensive PBL framework in medical education.

3
Specialty Choice Attitudes Among Medical Interns: Evidence from Hormozgan University of Medical Sciences

Kashefi Sis, P.; Shendabadi, A.; Alimi, N.; Boushehri, E.

2026-06-15 medical education 10.64898/2026.06.12.26355502 medRxiv
Top 0.1%
59.1%
Show abstract

Background: Choosing a medical specialty is a critical career decision that affects both physicians future professional lives and the composition of the healthcare workforce. Specialty preferences are shaped by multiple personal, educational, and socioeconomic factors, yet evidence from senior medical students in southern Iran remains limited. This study aimed to assess willingness to pursue specialty training among medical interns at Hormozgan University of Medical Sciences, identify their preferred specialties, and examine factors associated with their decisions. Methods: This descriptive-analytical cross-sectional study was conducted in 2023 among medical interns at Hormozgan University of Medical Sciences in Bandar Abbas, Iran. Using a convenience census approach, all eligible interns were invited to participate, and 83 students completed an online questionnaire. The instrument collected demographic, academic, and occupational data, as well as reasons for willingness or unwillingness to pursue specialty training and specialty preferences. Content and face validity were assessed by faculty members and students, and internal consistency reliability in the present study was acceptable (Cronbach alpha = 0.82). Data were analyzed using descriptive statistics and logistic regression in SPSS version 27. Results: Of the 83 participants, 50 (60.2%) reported willingness to pursue specialty training, while 33 (39.8%) did not. Among students willing to continue, the most frequently cited reasons were achieving a better economic position, broader job opportunities, and higher social status. Among those unwilling to continue, the most common reasons were fatigue from prolonged studying, financial problems, and the desire to start working after graduation. Radiology was the most common first-choice specialty, followed by otorhinolaryngology, dermatology, and cardiology. In regression analyses, no demographic or academic variable remained independently associated with willingness to pursue specialty training in the final multivariable model. Conclusions: A majority of medical interns were interested in pursuing specialty training, with preferences concentrated in a limited number of specialties perceived as offering favorable financial prospects, prestige, and lifestyle. Economic concerns and educational fatigue were the dominant factors influencing willingness and unwillingness to continue specialty education. These findings highlight the need for structured career counseling, broader exposure to different specialties, and policy measures to address financial and structural barriers to residency training. Keywords: medical specialty choice; medical interns; residency training; medical education; Hormozgan university of medical sciences

4
From Classroom to Clinical: A Mixed methods study of the Impact of Simulation Based Education on Final Year Physiotherapy Students

McGurn, C.; George, L.

2026-08-25 medical education 10.64898/2026.08.21.26360984 medRxiv
Top 0.1%
56.1%
Show abstract

Objectives: This study aimed to examine physiotherapy students reaction and learning, following Simulation Based Education (SBE) during a Cardiorespiratory module. A further aim was to see if any learning translated into clinical placement. Design: A mixed methods research design consisting of a questionnaire (Phase1) after the activity which was underpinned by the Kirkpatrick model of evaluation. This was followed by a focus groups (Phase 2) after completion of clinical placement. Participants: 92 final year physiotherapy students at a single institution were eligible to participate in the SBE session with n=81 (88%) students completing the survey, and 8 students participating across 2 focus group sessions. Results: Survey: Over 80% of students strongly agreed on a positive initial reaction to SBE. Learning yielded a 76% and above response of strongly agree in all areas except confidence. Students valued SBE as a preferred learning and teaching strategy and wanted more. They welcomed SBE as a supplement but not substitute for clinical placement. Students felt skills learned could be transferred into all areas of clinical practice,, namely communication and decision making. Conclusion: Students rated SBE positively with development of transferable non-technical skills. Reaction to SBE was high in terms of relevance, engagement and satisfaction. Self-perceived confidence, although positive, was the comparatively lowest scoring of the domains. Students advocated SBE as a supplement rather than a substitute for clinical placement preferring a hybrid approach. Contribution of the Paper: Adds to the positive body of evidence which already exists towards SBE, especially in the development of non-technical skills. Suggestions are made that SBE cant make students feel fully confident in preparation for clinical placement. While students in this study expressed that they wanted more SBE and earlier, this paper found that students would not advocate SBE replacing clinical placement.

5
From Prompt to Patient: "Cost-Effective Simulation Tool for Medical Education in Low-Resource Settings" Evidence from Mozambique

Impito, P. F.

2026-08-17 medical education 10.64898/2026.08.13.26360321 medRxiv
Top 0.1%
46.9%
Show abstract

Medical education in resource-limited settings faces significant challenges in providing diverse clinical exposure and fostering essential skills such as clinical reasoning, communication, and empathy. Due to the inability to afford immersive technologies such as virtual reality (VR) and Augmented Reality (AR), constrained by financial, infrastructural, and structural barriers, interactive simulation videos (ISV) constitute an innovative, cost-effective educational tool that can bridge the gap between theoretical knowledge and practical clinical experience, while enhancing student engagement and learning outcomes. This study aimed to assess the educational value and user experience of ISV as a supplementary tool in teaching medical semiology among medical students in Mozambique. A quantitative, descriptive, cross-sectional study was conducted among 4th-year medical students at Alberto Chipande University. Descriptive and inferential statistical analyses were performed, including a one-sample t-test to compare responses against a value. A total of 93 students participated in the study. The ISV was highly rated for realism (77.4%), relevance to training (74%), and usefulness of feedback (78.5%). Most students reported increased confidence in patient care (81.7) and found the digital patient credible and engaging (74.2). Overall mean scores across all domains were significantly higher than the neutral benchmark (p<0.05), indicating a positive perception of the tool. Students also expressed a strong willingness to recommend its integration into medical curricula. In conclusion, ISV represents a valuable and feasible pedagogical approach in medical education, particularly in low-resource settings. They enhance clinical reasoning, engagement, and confidence while providing scalable, standardized learning experiences. ISV holds strong potential as a complementary tool to bridge gaps in traditional medical training and improve educational equity.

6
Design, Implementation, and Evaluation of a Shadowing Program for Medical Students in the Basic Sciences Phase

Omid, A.; Changiz, T.; ghasemi, s.; Khodadoustan, z.; Heshmat, K.; Arefan, A.; Fazel Harandi, M. H.; Yousefi, M.

2026-06-12 health policy 10.64898/2026.06.10.26355363 medRxiv
Top 0.1%
45.2%
Show abstract

Introduction Shadowing, as an educational method based on active observation, can foster a realistic understanding of professional roles and enhance the communication skills of medical students. This study aimed to design, implement, and evaluate a shadowing program for basic sciences medical students. Methods This development study was conducted based on the ADDIE model in five phases. The study population consisted of 799 medical students in semesters 2 to 5. The stages included Analysis (determining needs through literature review and expert panels), Design (specifying learning environments and evaluation methods), Development (preparing guides and educational tools), Implementation (within the Medical Ethics course), and Evaluation (using questionnaires and reflection forms). Findings This study aimed to design and evaluate an educational shadowing program based on the ADDIE model. In the Analysis phase, the profiles of 799 students and learning objectives were determined. In the Design phase, a structured program for four types of shadowing was designed. In the Development phase, all guides and educational tools were prepared. In the Implementation phase, the program was carried out with complete coverage and adherence to ethical considerations. Finally, the program evaluation showed that "Motivation to become a good physician" (3.75-3.95) and "Enhancing empathy" (3.50-3.94) received the highest scores, while "Increasing understanding of the basic science-clinical connection" (2.53-2.89) and "Willingness to attend on holidays" (1.87-2.31) received the lowest scores. Conclusion The findings indicate that implementing the shadowing program is an effective method for strengthening the professional attitudes and academic motivation of medical students. However, the program did not significantly improve students perception of the basic science-clinical connection, indicating a need for curricular refinement. The continuation and extension of this program to other levels and fields of medical sciences are recommended.

7
Integrative Mechanisms of Early Clinical and Research Training (ECART) in Orthopaedic Medical Education: A Qualitative Single-Case Study

Lou, Y.; Liu, H.; Xu, X.; Xiao, Y.; Ma, D.; Shen, W.; Wang, C.; Kong, X.; Feng, S.

2026-06-12 medical education 10.64898/2026.06.11.26355438 medRxiv
Top 0.1%
42.9%
Show abstract

Background: Early clinical exposure and student participation in research are important components of medical training. They may support learning motivation, evidence literacy, and self-directed learning. In many programmes, however, clinical training and research training remain separated. Few studies have explained, within a real teaching team, how learners turn clinical phenomena into researchable questions and how research participation can reshape their clinical understanding. Early Clinical and Research Training (ECART) is a clinical-research integration approach developed by an orthopaedic team at the Second Hospital of Shandong University. Methods: We conducted a theory-informed, interpretivist qualitative single-case study. The case was an orthopaedic clinical-research team at the Second Hospital of Shandong University. Participants included medical undergraduates, academic degree graduate students, professional degree graduate students, clinical teachers, and research platform leads. We used purposive sampling with maximum variation. Data were collected through semi-structured interviews and de-identified teaching documents. Data were analysed using the framework method and were interpreted with a Context-Activity-Mechanism-Outcome (CAMO) logic. Results: The analysis showed that ECART was not simply early entry into the clinic or early entry into the laboratory. It was a team-based learning process centred on real medical problems. Four themes were identified. First, early clinical exposure helped learners make real problems visible and nameable, rather than merely increasing exposure. Second, clinical-research connection followed different pathways. Professional degree graduate students often started from clinical uncertainties in residency training and case management, and moved toward evidence-informed small projects. Academic degree graduate students often started from literature gaps, experimental findings, and mechanistic hypotheses, and then used clinical feedback to calibrate meaning. Third, research training, through literature reading, group meetings, experimental design, data review, and mentor questioning, helped learners move from completing tasks to explaining problems. Fourth, sustained ECART depended on a tiered team ecology formed by clinical teachers, research mentors, research platforms, and senior peers. Based on these findings, we refined the ECART programme theory: real medical problems are translated through explanation, searching, experimentalisation, and feedback-based reinterpretation into research questions that learners can understand, discuss, and test. This process supports problem formation, evidence awareness, mechanistic reasoning, translational judgement, and career clarification. Conclusion: ECART is best understood as a clinical-research integrated learning ecology that emerges from real team practice, rather than as a fixed standardised course. Its educational value lies in a recurring cycle of real problems, research translation, multi-source feedback, and clinical reinterpretation. This framework may inform the design, evaluation, and contextual adaptation of clinical-research integration pathways in medical education.

8
Validation of an Assessment Scale for a Low-Tech Laparoscopic Appendectomy Simulation and Its Relevance for Formative Self-Assessment

Tumameu Kouam, T. H.; Renoult, L.; Poitevin, M.; Jourdin, L.; Herve, C.; Meignan, P.; Podevin, G.; Schmitt, F.

2026-07-21 medical education 10.64898/2026.07.20.26358477 medRxiv
Top 0.1%
40.9%
Show abstract

Introduction: Laparoscopic appendectomy is an ideal procedure for acquiring laparoscopic skills through simulation. Nevertheless, technical training is time consuming for surgical trainers to provide constructive feedback, but this could be improved by the development of validated tools that enable appropriate formative self-assessment. For this reason, we developed a structured assessment scale for a laparoscopic appendectomy exercise using a low-fidelity simulator. The objective of this study was to validate the scale for use in formative self-assessment. Methods: During laparoscopic simulation sessions in 2025-2026, participants with varying levels of experience performed a standardized laparoscopic appendectomy (LAP) exercise on a low-fidelity simulator. Performance was assessed through formative self- and external assessment using a specific scale derived from the OSATS (Objective Structured Assessment of Technical Skills) score. Content and construct validity, internal consistency, reproducibility, and reliability in both hetero- and self-assessment were analyzed. Results: Thirty-two participants were included in the validation study of the LAP scale, including 7 medical students, 17 residents in pediatric, visceral, urological, and gynecological surgery, and 8 practicing surgeons. The content of the scale was deemed relevant by 80% of the users. It demonstrated excellent construct validity, with scores increasing according to level of experience: 9.9 +/- 0.7 among students, 12.7 +/- 3.3 among junior residents, 16.6 +/- 3.3 among experienced residents, and 18.8 +/- 0.9 among practicing surgeons (p < 0.0001). Reproducibility and internal consistency were significant, while inter and intrarater reliability were excellent (correlation coefficients r = 0.90 and 0.91; p < 0.0001), as was the correlation between external and self-assessment (r = 0.81; p < 0.0001). Self-assessment was more reliable among experienced learners than among novices. Conclusion: This standardized LAP scale is validated for both external and self-assessment, the latter requiring prior training to be reliable and formative.

9
High Demand, Low Possession: Dilemmas and Strategies for Research Capability Cultivation in Clinical Medicine Postgraduates

Wang, B.; Yang, L.; Gong, Z.

2026-06-15 medical education 10.64898/2026.06.12.26355561 medRxiv
Top 0.1%
40.5%
Show abstract

Most previous studies have examined medical postgraduate research training from a single dimension, lacking a full-chain analysis that integrates capability demand, actual possession, obstacles, and output. Consequently, the measurement of capability gaps and the analysis of underlying training model deficiencies remain insufficient. To address this gap, we administered a self-designed multidimensional questionnaire to 86 clinical medicine postgraduates at a medical school, covering research cognition, interest, capability demand and possession, participation pathways, difficulties, and outputs. The aim was to systematically characterize the current situation, identify problems, and propose optimization strategies. Over 90% of participants expressed interest in research, yet only 1.16% self-rated as very knowledgeable. The largest demand-possess gap was for writing and publication (86.05% vs. 16.28%), followed by independent research capability (75.58% vs. 11.63%). A total of 59.30% cited lack of foundational knowledge, making experiments very difficult, as the greatest challenge, and 66.28% had no research achievements. The primary source of research topics was supervisor assignment (54.65%), with only 4.65% choosing topics independently. No statistically significant differences were found across grades or training types (P > 0.05). These findings reveal a structural high demand, low possession gap in medical postgraduate research training, with early research experience deficit and a passive research model as key constraining factors. Accordingly, an integrated bachelor-postgraduate progressive research competency training system is proposed.

10
Diagnostic analytics of routine Clinical Competency Committee data of six cohorts in family medicine program in the UAE, utilizing Milestones, EPA, and ITE

Baynouna Alketbi, L. M.; Nagelkerke, N.; Alzarouni, A.; AlKwuiti, M.

2026-07-16 medical education 10.64898/2026.07.13.26356644 medRxiv
Top 0.1%
36.0%
Show abstract

In Competency-based medical education (CBME), longitudinal data is generated continuously. The judgments a Clinical Competency Committee (CCC) makes about trainee learning and performance are a valuable resource, supporting both resident and program development. Such data as well can enables the evaluation of rating quality and of CBME instruments such as Milestones and Entrustable Professional Activities (EPAs) which can help address a gap in the CBME literature, where evidence on the performance of these instruments remains limited. Objective Routinely gathered CCC data of six cohorts in a four-training ACGME-I-accredited family medicine residency in Al Ain, United Arab Emirates, was studied to describe growth trajectories, rating-system behavior, and the concurrent agreement of CBME instruments. As well as investigating the prospective predictive validity of two CBME instruments, EPA and Milestones, and the In-Training Exam (ITE). Methods The longitudinal CCC data for 80 residents across six cohorts (2019-20 to 2024-25) were assessed at up to eight time points (mid- and end-year; R1-R4). The pooled dataset included 10,458 EPA item ratings across 334 resident-time points, 5,021 Competency Milestone item ratings across 285 resident-time points, and 185 ITE scores. Five research questions were examined: growth trajectories; within- and between-resident variation and straight-lining (identical scores on assessed items at a single time point); EPA-Milestones agreement; the validity of supervisor ratings against the ITE (anchoring diagnostic, same-year correlations, prospective regressions); and EPA blueprint fidelity (the mapping of EPAs against the ACGME-I subcompetency). Al Ain trajectories were benchmarked against an international family medicine reference. Results All three instruments rose steadily across the eight timepoints. By End R4, the Milestones mean (4.00, range 3.83-4.24) matched US end-of-training norms (3.84-4.02). With regards to rating quality, pooled R1-R3 Milestones straight-lining was 2.3% (EPA 0%), below US benchmarks; between-resident discrimination was preserved (SD 0.41-0.54); and longitudinal halo was ruled out (within-domain growth-slope r = 0.61 vs across-domain r = 0.37). End R1 Overall EPA was the strongest prospective predictor of Final Competency (B = 0.96, p < 0.001) and Final ITE (B = 96.88, p = .006). Medical Knowledge ratings were independent of prior ITE scores from Mid R2 onward, and End R2 MK ratings predicted ITE 17 months later at r=0.88, confirming supervisor judgment was not anchored to test results. With regards whether individual EPAs correlate with individual Milestone subcompetencies at each timepoint, a significant EPA and Milestones correlations were negligible at End R1 (1 of 222 item-level cells significant) and converged by End R3 (36 cells), while resident-mean stepwise regressions showed the two instruments (EPA and Milestones) behaved as overlapping predictors throughout, indicating that EPAs and Milestones are complementary at the level of specific content but convergent at the level of aggregate resident judgment. Blueprint fidelity rose from 30% of cells reaching r [&ge;] 0.40 at End R2 to 80% at End R3 in the same cohort, indicating that apparent fidelity is materially affected by measurement timing. Conclusion By graduation, residents demonstrated substantial and progressive competency achievement across both instruments, with the majority reaching the entrustable threshold on both EPA and Milestone ratings. The rating system demonstrated disciplined assessment behavior of supervisors and both concurrent and prospective validity relative to the ITE. Overall EPA at End R1 was the strongest prospective predictor of all three terminal outcomes, final ITE score, graduating Competency Milestones, and graduating overall EPA, outperforming Milestones and baseline knowledge. Routine CCC data support an evidence-based quality assurance framework spanning rater-process diagnostics, outcome-validity diagnostics, and the asymmetric-instrument diagnostic, requiring no additional data collection beyond existing program processes.

11
Students' Perceptions of an AI-Enhanced Ethics Learning Platform: A Pilot Study on Interprofessional Healthcare Education

Rankine, L.; Van Bussel, J.; Moodie, S. T.; Tawiah, A. K.

2026-06-26 medical education 10.64898/2026.06.23.26356394 medRxiv
Top 0.1%
35.5%
Show abstract

Introduction: Generative artificial intelligence (AI) can produce realistic clinical scenarios on demand and deliver immediate, individualized feedback, yet its use to teach ethical reasoning, rather than to address the ethics of AI itself, remains underexplored in interprofessional healthcare education. Aim: This pilot study examined how interprofessional healthcare students perceived an AI-enhanced, case-based platform designed to support ethical decision-making across physical therapy, occupational therapy, speech-language pathology, and audiology. Methods: Students enrolled in an interprofessional education course completed an online module of 20 instructor-vetted, AI-generated ethics cases and an optional post-activity survey of Likert-scale and open-ended items. Quantitative data were analyzed descriptively and qualitative responses were analyzed through content analysis. Results: Ten students responded. Within this small sample, perceptions of platform utility and usability were strongly positive, with all respondents agreeing that immediate feedback and scenario variety supported learning. Perceptions were more divided when the platform was compared directly with traditional classroom learning, and respondents identified pacing and auto-scrolling as usability concerns. Conclusions: These preliminary findings suggest AI-enhanced case-based platforms can engage students and support applied ethics learning but are best positioned to complement rather than replace traditional instruction. Findings are exploratory given the small, demographically limited sample.

12
Where Do I Belong? Searching for fit in an unseen specialty: medical students paths to Youth Health Care

Muyselaar-Jellema, J. Z.; Könings, K. D.; van Dijk, A.; Kiefte-de Jong, J. C.; Nierkens, V.

2026-07-16 medical education 10.64898/2026.07.14.26358057 medRxiv
Top 0.1%
31.2%
Show abstract

Introduction: As healthcare systems increasingly shift toward prevention and community-based care, the demand for physicians in extramural specialties continues to grow. Yet, a mismatch persists between workforce needs and medical students career aspirations. Little is known about how medical students and trainees develop an interest in extramural specialties such as youth health care (YHC). This study explores trainees trajectories toward becoming a youth health care physician (YHCP). Methods: We conducted a qualitative study using semi-structured online interviews with fourteen YHCPs in training. We combined an inductive and deductive approach, applying the person-environment (PE) fit framework to explore participants evolving experiences of fit and misfit. Results: Participants described growing misfit with clinical culture of medical school (i.e. the hidden curriculum), particularly during hospital-based clerkships, combined with limited exposure to YHC. For some, this misfit extended to doubts about becoming a doctor. Over time, participants developed a sense of fit and belonging within YHC, either directly or after exploring other specialties including extramural specialties. Discussion: These findings reframe specialty choice as a longitudinal search for belonging and alignment, in which trainees iteratively explore, evaluate, and refine their sense of fit across contexts. Clerkships serve as key sites for testing fit, yet also expose learners to the clinical culture, including the hidden curriculum. Broadening exposure and supporting reflective fit processes may encourage more medical students to choose extramural specialties, ultimately fostering a more balanced and sustainable alignment of the medical workforce.

13
What level of expertise is necessary to generate ACLS training test questions: pre-med students vs. artificial intelligence?

LoGalbo, S. S.; Richman, M.; Wang, J.; Saji, I.; Traore, A.; Oliva, H.; Wu, E.; Drudi, A.; Foster, D.; Bhandari, S.; Delfillo, R. L.; McCann, A.; Coard, J.; Matthew, C.; Smith, B.

2026-06-11 medical education 10.64898/2026.06.11.26354470 medRxiv
Top 0.1%
26.8%
Show abstract

Abstract Introduction In-hospital cardiac arrest carries high mortality despite standardized ACLS training. Educators face increasing time constraints in developing assessment tools for ACLS training. Two possible solutions to this problem are using pre-medical students or using artificial intelligence to generate test questions. This study compared the quality of pre-medical student-generated ACLS test questions vs. AI-generated ACLS test questions, testing the hypothesis that AI-generated questions are non-inferior to student-generated questions. Methods Ten pre-medical students created ACLS questions following predefined criteria, while an AI model (Northwell's Artificial Intelligence Hub) generated comparable questions. A blinded ACLS-certified physician evaluated questions on the qualities of Alignment, Clarity, Cognitive Level, and Question Design using a standardized rubric (Likert scale: 1 = poor quality, 5 = excellent). Student's T-test and Chi-square analysis were used to compare the quality of questions on different rubric domains within each arm (student vs. AI) and within one domain (eg, question Clarity) between arms. The Student's T test was used when 2 comparator groups were compared (eg, Clarity of student-generated vs. AI-generated questions) within one arm. The ANOVA test was used when comparing more than 2 comparator groups (eg, Alignment vs. Clarity vs. Cognitive Level) within one arm. Statistical significance was set as a priority at p <0.05. Results Both student-generated and AI-generated questions were of high quality. AI-generated questions achieved the maximum score in the domains of Alignment, Clarity, and Question Design, but fell short of perfect scores in the domain of Cognitive Level (8 of 50 questions were less than 5). Student-generated questions achieved less-than-perfect scores in each domain. No significant difference was found in overall mean question scores between groups (students = 4.79, AI = 4.81; p = 0.9). However, AI-generated questions had significantly-greater Clarity (students = 4.8, AI = 5; p = .0461), while Alignment, Cognitive level, and Question Design showed no significant differences. Conclusion AI-generated questions demonstrated overall quality comparable to those generated by pre-medical students, supporting the potential role of AI as a scalable tool in ACLS educational assessment development. Further studies are warranted to evaluate additional AI platforms and determine optimal integration of AI in medical education assessment design.

14
Establishment and Efficacy of an Endoscopic Pathogen Visualization Literacy (EPVL) Training Program for Gastroenterologists Based on Fluorescence Rapid On-Site Evaluation (ROSE) Technology

Zhang, L.; Hou, Y.; Li, B.; Wu, K.; Zhang, j.; Yang, M.

2026-08-13 medical education 10.64898/2026.08.12.26360123 medRxiv
Top 0.1%
24.0%
Show abstract

ObjectiveTo establish a standardized training program for endoscopic pathogen visualization literacy (EPVL) based on fluorescence rapid on-site evaluation (ROSE) technology for gastroenterologists, and to evaluate its training efficacy. MethodsA prospective quasi-experimental study was conducted. A total of 54 gastroenterology trainees were non-randomly allocated into the EPVL training group (Group A, n=28, 16-hour comprehensive training) and the control group (Group B, n=26, 3.5-hour traditional teaching). Pre- and post-training assessments included theoretical examinations, fluorescence ROSE image interpretation tests (30 parallel images per set), interpretation speed measurement, and clinical decision-making integration evaluation. The primary outcome was the change in image interpretation accuracy, analyzed by ANCOVA with pre-test scores as the covariate. ResultsBaseline characteristics were comparable between groups (P>0.05 for all demographic variables and pre-test scores). Group A showed significant improvement in image interpretation accuracy from 57.8{+/-}13.6% pre-training to 82.5{+/-}11.2% post-training (improvement of 24.7%, paired t=-12.86, P<0.001), while Group B improved from 58.5{+/-}13.0% to 71.0{+/-}13.5% (improvement of 12.5%, paired t=-5.24, P<0.001). After ANCOVA adjustment for pre-test scores, the between-group difference was significant (F(1, 51)=10.95, P=0.0017, 2=0.177), with Cohens d=0.94 (large effect size). Interpretation speed in Group A (19.2{+/-}2.8 s/image) was significantly faster than in Group B (32.5{+/-}6.0 s/image, t=-10.45, P<0.001). Clinical decision-making scores were significantly higher in Group A (80.5{+/-}8.0 vs. 65.3{+/-}11.5, t=5.60, P<0.001). The Kappa agreement with the gold standard in Group A improved from 0.56{+/-}0.18 to 0.84{+/-}0.11 (t=-8.35, P<0.001). Participant satisfaction exceeded 88%. ConclusionThe EPVL training program significantly improves gastroenterologists fluorescence ROSE image interpretation accuracy, speed, and clinical decision-making integration, providing a novel and effective standardized training paradigm for digestive endoscopy education.

15
Stakeholder perspectives on implementing a maternal health blended learning course for health care providers in Kenya, Nigeria and Tanzania

Ladur, A. N.; Egere, U.; Murray, C.; Eyinda, M.; Muchemi, O. M.; Suleiman, Z.; Mdoe, M.; Rweyemamu, M. A.; Bello, A. B.; Mohammed, H.; Ameh, C. A.

2026-07-31 medical education 10.64898/2026.07.29.26359229 medRxiv
Top 0.1%
19.3%
Show abstract

Background Good quality maternity care is critical in reducing maternal morbidity and mortality in regions with a high maternal and perinatal mortality. Building capacity of maternity care providers through in service training has been proven as effective in bridging the knowledge, competency, and skills gap during provision of maternity care. Meaningful involvement of stakeholder perspectives in the design and implementation of public health training programs heightens the prospects of achieving long term changes in practice and policy. Methods To explore stakeholder perspectives and experiences on the co-development of the antenatal-postnatal care course in Kenya, Nigeria, Tanzania. Data was collected through nine key informant interviews and observation notes between July - October 2022. Qualitative data was analysed in NVIVO software using inductive thematic analysis. Results Study findings showed that stakeholders were receptive of the blended learning approach for training in antenatal and postnatal care describing it as accessible, and effective in improving knowledge and skills. Course facilitators and healthcare providers valued the interactive format, opportunities for pre-session preparation, and continuous support, which contributed to improved learning experiences and collaboration. Conclusion This study highlighted the importance of involving stakeholders in the design and implementation of the antenatal-postnatal blended learning course in three countries. Eliciting feedback based on stakeholder experiences and incorporating it in real time strengthened the implementation of the blended learning course.

16
A global cross-sectional survey of health professionals' interest-confidence gaps in value-based health care implementation: a learning needs assessment

Lewis, S.; Andrews, A.; Laing, H.

2026-06-11 medical education 10.64898/2026.06.10.26355253 medRxiv
Top 0.1%
19.2%
Show abstract

Abstract Objectives Value-Based Health Care (VBHC) increasingly guides health system redesign internationally. Despite the increasing availability of VBHC education, gaps remain between health professionals' conceptual understanding of VBHC and their confidence to implement it in practice. This study assessed perceived learning needs and preferences of healthcare professionals across foundational topics essential to VBHC implementation. Design Cross-sectional online survey study Setting and participants The survey was distributed to the global VBHC community and yielded 518 responses. Most respondents were based in the UK and Ireland (51%) and 65% had more than 10 years of experience in the health sector. Participants represented a variety of professional backgrounds, including clinicians (34%), operational or executive managers and leaders (22%), and life sciences or procurement professionals (13%). Primary and secondary outcome measures Primary outcome measures included self-reported interest and confidence across 15 VBHC domains and the magnitude of the gap between them. Secondary outcomes included perceived implementation challenges and preferred VBHC learning approaches, including prior engagement with VBHC-related learning. Results Respondents identified substantial VBHC implementation challenges, including implementing outcome measurement (62.4%), conflicting priorities (57.7%), and resistance to change (56.8%). Interest in all VBHC domains was high (median >= 80/10), while confidence to implement remained substantially lower across most domains (median <=50/100). The largest interest-confidence gaps were observed for reimbursement mechanisms, costing methodology, and overcoming implementation challenges. Interactive learning approaches, including in-person seminars/workshops (55.2%) and online masterclasses (53.9%) were preferred over self-directed formats. Conclusions This international survey identified consistent gaps between health professionals' interest in VBHC and their confidence to implement key VBHC domains in practice. Addressing these gaps through advanced, targeted and contextual education may support more effective and sustainable VBHC implementation in practice.

17
Characterizing large language model generative artificial intelligence variability in the production of objective structured clinical examination stations

Joseph-Delaffon, K.; Desgrouas, M.; Catanese, S.; Lejeune, J.; Nait-Kaci, J.; Piver, E.; Breteau, I.; Leducq, S.; Gatault, P.; Khanna, R. K.; Angoulvant, D.; Vallet, N.

2026-08-06 medical education 10.64898/2026.08.04.26359691 medRxiv
Top 0.1%
19.0%
Show abstract

Background. Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intelligence (AI) represents a promising path to accelerate content creation by automating the generation of scenarios. A growing number of AI tools is now available for this purpose. Objective. To assess the variability between generative AI models in their ability to produce OSCE stations in the field of paediatrics. Methods. A structured prompt was developed based on the French national OSCE guidelines for medical education. Five distinct AI models were provided with this prompt, alongside the neonatal jaundice chapter from the French pediatric reference textbook, to generate 6 complete OSCE stations. Results. Prompt compliance was high for ChatGPT 5.1, ChatGPT 5.2, Gemini 3.0 Pro, and Claude Opus 4.5, while it was lower for Grok 4.1. Expert-rated quality was generally high, with few factual errors or missing information across models. However usability differed significantly between models. This was also true for several quality dimensions such as checklist clarity, embedding of checklist answers within vignettes, and ease of standardized patient formation. ChatGPT 5.1 required the most revisions and Gemini most often rated usable as is. Significant inter-model differences were observed in diagnostics, only with ChatGPT 5.1 sampling all three neonatal jaundice categories. Contextual variables showed systematic narrowing across models. Clinical grid density was consistent (10-12 items per station), but thematic distribution differed markedly. Soft skills coverage varied significantly across models (p=0.002), none of them consistently representing all communication competency domains. Conclusion. Large language models can generate structurally compliant OSCE stations, but surface compliance conceals substantive inter-model differences in diagnostic coverage, contextual diversity, and soft skills representation, that compromise content validity. No model currently meets the criteria for unsupervised deployment in a summative assessment bank. The choice of model carries pedagogical implications and expert curation remains essential before integration into high-stakes assessment workflows.

18
AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills

Otte, J. H.; Cartagena, A.

2026-08-31 medical education 10.64898/2026.08.26.26361437 medRxiv
Top 0.1%
17.3%
Show abstract

Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.

19
Prospective study on the organization and efficiency of online journal club

Burlov, N.; Baranovskii, M.; Burlova, E.; Slavenko, M.; Khrykov, G.

2026-08-12 medical education 10.64898/2026.08.11.26360192 medRxiv
Top 0.1%
15.9%
Show abstract

Background. Journal clubs (JCs) are a popular education format. Interest in studying their impact is high, and authors often report positive results related to subjective parameters. Objective assessments of effectiveness are limited and contradictory. In this paper, we share our experience and describe our journal club effectiveness. Methods. We conducted a prospective cohort study within our online journal club. Meetings followed a discussion-based format and were held via Zoom, with timing and topics determined by voting in the club Telegram chat. Enrolment occurred in waves and included an application, entry test, and interview. During each recruitment wave, both club members (treatment group) and applicants (control) completed an admission test assessing knowledge of evidence-based medicine and statistics. Results. The JC currently comprises 27 members. Over the past year, 76 meetings were held, with 75% of participants grading their experience with 9 or 10 on a ten-point scale. Multivariate analysis demonstrated non-significantly results (SMD = 0.19 (95% CI 0.004; 0.38), p = 0.046) among participants. However, in other adjusted models, differences between groups were not statistically significant (p > 0.05). Conclusion. While the analysis of the subjective outcomes is consistent with findings from previous studies, the objective outcomes remain inconclusive. Further research is needed to refine the methodology for the organization and evaluation of journal clubs.

20
VICTORY Protocol - VIrtual knowledge exchange in primary Care Through effective digital Online couRses for all Young people without borders and barriers

Brady Bates, O.; Mbakaya, B.; Cullen, W.; Gallagher, J.

2026-07-14 medical education 10.64898/2026.07.11.26357803 medRxiv
Top 0.1%
15.5%
Show abstract

Aim This study aims to design, implement, and evaluate a virtual exchange programme in global health and primary care, with a focus on building capacity and fostering collaboration between European and Sub-Saharan African institutions. Methods A mixed-methods approach will be adopted across three phases. First, a scoping review will synthesise evidence on the implementation, outcomes, barriers, and facilitators of virtual exchanges in global health. Second, a virtual exchange curriculum will be co-developed through surveys and focus groups with medical students and educators. Third, the programme will be evaluated using pre- and post-course surveys, and semi-structured interviews to measure changes in knowledge, global health competencies, and intercultural awareness. Quantitative data will be analysed using descriptive and regression analyses, while qualitative data will undergo reflexive thematic analysis. Conclusion This study will generate evidence on the design and impact of virtual exchanges for global health education, contributing to the development of sustainable and equitable curricula in PHC. By disseminating findings across academic, policy, and professional networks, the VICTORY project seeks to advance global collaboration, support workforce development, and promote health equity.